Papers by Tiago Timponi Torrent

6 papers
Domain Adaptation in Neural Machine Translation using a Qualia-Enriched FrameNet (2022.lrec-1)

Copied to clipboard

Challenge: Neural models have been advancing in a myriad of tasks, but there is a lack of large training data.
Approach: They propose a method for domain adaptation of Neural Machine Translation systems using a multilingual FrameNet enriched with qualia relations as an external knowledge base.
Outcome: The proposed system outperforms the state-of-the-art commercial system in an experiment . the proposed system substitutes domain-specific terms in the source language by their adequate translation in the target language.
Framed Multi30K: A Frame-Based Multimodal-Multilingual Dataset (2024.lrec-main)

Copied to clipboard

Challenge: Recent advances in image-captioning datasets combine image and language to solve a diverse range of tasks.
Approach: They propose a Brazilian Portuguese multimodal-multilingual dataset that extends the Multi30K dataset with 158,915 original Brazilian Portuguese descriptions and 30,104 Brazilian Portuguese translations.
Outcome: The proposed dataset adds 2,677,613 frame evocation labels to the 158,915 English descriptions and to the ones created for Brazilian Portuguese.
Frame Shift Prediction (2022.lrec-1)

Copied to clipboard

Challenge: Frame shift is a cross-linguistic phenomenon in translation which results in corresponding pairs of linguistic material evoking different frames.
Approach: They propose a task to predict cross-linguistic frame-to-frame correspondence and propose auxiliary training to learn cross-lingual frame-by-frame correlation.
Outcome: The proposed task can learn cross-linguistic frame-to-frame correspondence and predict frame shifts in a Berkeley FrameNet-like configuration.
Frame2: A FrameNet-based Multimodal Dataset for Tackling Text-image Interactions in Video (2024.lrec-main)

Copied to clipboard

Challenge: et al., 2016) describe a multimodal dataset built from a Brazilian travel TV show . frameNet is composed of frames and their associated roles in a network of typed frame-to-frame relations.
Approach: They present a multimodal dataset built from a Brazilian travel TV show annotated for FrameNet categories for both text and image communicative modes.
Outcome: The proposed dataset includes 230 minutes of video annotated for FrameNet categories . the model can be applied to other communicative modes, i.e., images .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations